# MOLECULAR OLFACTION ARCHITECTURE (MOA) ### A Conceptual Framework for Olfactory Perception in Large Language Models **Harpia AI Research** *Concept Paper — Not peer-reviewed. Presented as a speculative technical proposal.*
──────────────────────────────────────────────────────────────────────────────── ABSTRACT Large language models (LLMs) have achieved multimodal perception across vision, audio, and text. Olfaction — the sense of smell — remains one of the major human sensory modalities without a corresponding digital input modality for LLMs. This paper proposes the Molecular Olfaction Architecture (MOA), a conceptual framework in which a molecular detection layer identifies volatile organic compounds (VOCs) present in an environment and passes them as structured input to an LLM. The LLM then applies its learned chemical and semantic knowledge to produce a natural-language interpretation of the detected scent. We hypothesize that LLMs already possess substantial implicit knowledge of chemistry and olfaction acquired during pretraining, and that MOA could provide the missing sensory bridge between physical molecular detection and semantic reasoning. An informal proof-of-concept demonstrates the potential viability of the reasoning layer independently of physical sensing hardware. This paper describes the proposed architecture, its potential applications, limitations, and directions for future empirical validation. 1. INTRODUCTION The development of multimodal artificial intelligence has followed a relatively consistent pattern: connect a perception encoder to a language model, and the model gains the ability to reason about information originating from a new sensory modality. Vision-language models such as GPT-4V and Gemini demonstrate this paradigm for visual information. Audio-language models extend similar capabilities to speech and environmental sound. In these systems, the underlying language model does not necessarily need to be fundamentally redesigned to process a new modality. Instead, it can receive a structured representation of sensory information through an appropriate input interface. Olfaction has not yet followed the same path. Despite being one of the most chemically complex and information-rich human senses, smell does not currently have a standardized digital input modality integrated into general-purpose LLM systems. Current LLMs can discuss smells and describe the expected odor of substances such as coffee, rain, gasoline, or flowers. However, they cannot directly perceive these odors from the physical environment. The missing component is an olfactory encoder capable of converting molecular information into a representation that an LLM can process. This paper proposes the Molecular Olfaction Architecture (MOA) as a conceptual solution. MOA consists of a molecular detection layer that identifies volatile compounds present in the surrounding environment and an LLM reasoning layer that interprets those compounds semantically, producing a human-readable description of the corresponding olfactory profile. The central hypothesis is that the semantic knowledge required for this interpretation may already exist within general-purpose LLMs. If this is the case, the primary missing component is not necessarily a new model architecture, but rather a reliable sensory input channel. 2. BACKGROUND 2.1 ELECTRONIC NOSE TECHNOLOGY Electronic nose (e-nose) technology has existed for decades. Traditional e-noses typically employ arrays of chemical sensors, including metal oxide semiconductor (MOS) sensors and conductive polymer sensors, to generate an electrical fingerprint associated with an odor. These fingerprints are commonly processed using pattern-recognition and machine-learning techniques such as principal component analysis (PCA), support vector machines (SVMs), and other classification methods. Although effective for specific applications, this approach is fundamentally limited by the scope of its learned classification space. An e-nose generally identifies odors according to previously defined classes or reference samples and produces a predefined classification output. Such systems are not inherently designed to perform open-ended semantic reasoning about molecular compositions. For example, an e-nose may identify a sample as "coffee," but it does not necessarily possess the ability to reason about why the sample smells like coffee, which compounds contribute to the perception, or what a novel combination of compounds might imply. LLMs provide a potentially complementary capability. 2.2 LLMS AND CHEMICAL KNOWLEDGE Through pretraining on scientific literature, chemical databases, technical documentation, and general text, LLMs may acquire associations between chemical compounds and their known properties, including olfactory characteristics. For example, an LLM may associate geosmin with the characteristic earthy odor commonly perceived after rain on dry soil, or associate 2-furfurylthiol with roasted coffee aroma. This knowledge is normally latent and cannot be directly triggered by real-world molecular measurements because current LLM systems generally lack an olfactory sensory interface. MOA proposes connecting these two domains: MOLECULAR SENSING │ ▼ CHEMICAL INFORMATION │ ▼ STRUCTURED REPRESENTATION │ ▼ LLM │ ▼ SEMANTIC INTERPRETATION │ ▼ NATURAL-LANGUAGE OLFACTORY OUTPUT 3. THE MOA ARCHITECTURE MOA proposes a three-stage processing pipeline. 3.1 STAGE 1 — MOLECULAR DETECTION A chemical sensor array, electronic nose, gas chromatography system, or mass spectrometer samples the ambient air and estimates which volatile organic compounds (VOCs) are present. Depending on the sensing technology, the output may include: • Compound identities • Confidence values • Estimated concentrations • Relative abundance • Detection timestamps • Sensor reliability This stage is analogous to an image encoder in a vision-language system. It converts a physical phenomenon into structured digital information that can subsequently be processed by an AI model. 3.2 STAGE 2 — STRUCTURED INPUT FORMATTING The detected compounds are transformed into a standardized representation suitable for LLM processing. A simplified example could be: 2-Furfurylthiol: high Pyrazines: high Diacetyl: medium Guaiacol: low Acetic acid: low A more advanced representation could include estimated concentrations, confidence scores, molecular identifiers, sensor reliability, and environmental metadata such as temperature and humidity. For example: Compound: 2-Furfurylthiol Estimated concentration: high Detection confidence: 0.94 Compound: Pyrazines Estimated concentration: high Detection confidence: 0.88 Compound: Diacetyl Estimated concentration: medium Detection confidence: 0.81 3.3 STAGE 3 — LLM SEMANTIC REASONING The structured molecular representation is provided to an LLM as sensory input. The LLM applies its learned knowledge of chemistry, molecular associations, odor descriptors, and environmental context to infer a likely olfactory profile. The resulting output may include: • A predicted scent or combination of scents • A natural-language description of the odor • The likely contribution of individual compounds • Confidence estimates • Possible environmental sources • Relevant contextual interpretations • Identification of unusual molecular patterns The key architectural hypothesis is that Stage 3 may not require a specialized olfactory language model or extensive fine-tuning. If general-purpose LLMs already contain sufficient chemical and olfactory knowledge, MOA primarily needs to provide a reliable sensory input channel capable of exposing that knowledge to real-world molecular data. 4. INFORMAL PROOF OF CONCEPT To evaluate the reasoning layer of MOA independently of physical sensing hardware, an informal proof-of-concept test was conducted. A list of volatile compounds associated with freshly brewed roasted coffee was manually composed and provided to a general-purpose LLM (Google Gemini). The prompt followed this structure: "You are a test of a new architecture emerging for olfaction in LLMs. Identify this scent and I will tell you if you are correct: 2-Furfurylthiol, Geosmin, Diacetyl, Pyrazines, Acetic acid, Formic acid, Guaiacol, Furaneol." The model identified the target scent as freshly brewed roasted coffee and provided a detailed interpretation of the potential contribution of the listed compounds to the overall olfactory profile. 4.1 CRITICAL LIMITATION This experiment has a critical limitation: The molecular input was manually constructed by a human who already knew the target scent. The compound list was not generated by an independent molecular sensor or spectrometry system. Therefore, this experiment does not constitute empirical validation of the complete MOA pipeline. It should instead be interpreted as a preliminary demonstration that the semantic reasoning layer can accept a molecular representation and generate a plausible olfactory interpretation using an existing general-purpose LLM without architectural modification or task-specific fine-tuning. A rigorous evaluation would require: • Independently measured molecular samples • Controlled concentrations • Blind testing • Multiple scent classes • Comparison against human olfactory assessments • Comparison against established chemical reference data 5. POTENTIAL APPLICATIONS If implemented with sufficiently accurate and portable molecular sensing hardware, MOA could enable a class of applications that are currently difficult or impossible for conventional AI systems. 5.1 ROBOTICS Autonomous robots equipped with MOA could detect environmental chemical signatures and use them as an additional source of contextual information. Potential applications include: • Detection of gas leaks • Detection of smoke or combustion products • Identification of chemical spills • Food-quality assessment • Monitoring cooking processes through odor • Environmental monitoring • Search-and-rescue applications Instead of simply detecting a predefined chemical signature, the system could potentially reason about combinations of compounds and describe their meaning in natural language. 5.2 MEDICAL DIAGNOSTICS Human breath, skin emissions, and other biological samples contain volatile organic compounds that may correlate with physiological or pathological states. MOA could potentially assist researchers and clinicians by transforming detected volatile profiles into interpretable descriptions or hypotheses. Potential research applications could include: • Metabolic disorders • Infections • Certain cancers • Other conditions associated with measurable volatile profiles Such applications would require extensive clinical validation and should not be interpreted as established diagnostic capabilities. 5.3 FOOD AND BEVERAGE INDUSTRY MOA could enable real-time chemical and sensory quality monitoring. Rather than producing only a binary classification: PASS / FAIL a system could generate a semantic description such as: "Detected profile is consistent with roasted coffee, with elevated sulfur-containing volatiles and pyrazines." This could support: • Quality control • Production monitoring • Anomaly detection • Product consistency analysis 5.4 ACCESSIBILITY An olfactory interface could potentially provide individuals with anosmia or reduced olfactory perception with a digital representation of environmental smells. Instead of directly reproducing the physical sensation of smell, the system could translate molecular measurements into natural-language descriptions such as: Freshly cut grass. Strong citrus odor with a dominant lemon-like profile. Possible smoke detected. This could provide an alternative form of digital olfactory awareness. 6. LIMITATIONS AND OPEN PROBLEMS MOA faces several significant challenges that must be addressed before the architecture can be empirically validated. 6.1 HARDWARE COST AND ACCESSIBILITY Mass spectrometers capable of identifying specific molecular compounds at low concentrations can be expensive and difficult to miniaturize. Consumer-grade MOS sensor arrays are substantially more accessible, but many detect broad chemical responses rather than uniquely identifying individual molecules. This creates a fundamental trade-off between: Cost ↔ Portability ↔ Molecular Specificity ↔ Detection Accuracy 6.2 SENSOR DRIFT Chemical sensors can exhibit drift over time due to: • Aging • Environmental conditions • Contamination • Changes in sensor characteristics Long-term deployments would therefore require calibration procedures and methods for detecting and compensating for sensor degradation. 6.3 COMPOUND COMPLEXITY Real-world odors can consist of dozens or even hundreds of volatile compounds simultaneously. It remains unclear how accurately an LLM can interpret increasingly complex molecular mixtures, particularly when: • Compounds interact perceptually • Concentrations vary significantly • Mixture effects are nonlinear • The combination is poorly represented in training data 6.4 CONCENTRATION AND PERCEPTION The mere presence of a compound does not necessarily determine its perceptual importance. Human olfaction is influenced by: • Concentration • Odor thresholds • Molecular interactions • Mixture effects • Individual differences in perception Therefore, a future MOA system would likely need more than a simple binary list of detected compounds. 6.5 LLM HALLUCINATION RISK LLMs may produce confident but incorrect interpretations, particularly when presented with: • Unusual molecular combinations • Synthetic compounds • Incomplete measurements • Conflicting molecular profiles An operational MOA system would therefore require mechanisms such as: • Uncertainty estimation • Chemical verification • Retrieval-augmented generation • Structured validation • Sensor confidence propagation 6.6 NO END-TO-END EMPIRICAL VALIDATION The proof of concept described in Section 4 evaluates only the semantic reasoning component using manually constructed input. No complete end-to-end experiment involving real-time molecular sensing, automated compound identification, structured input generation, and blind scent identification has been conducted. Consequently, the feasibility of the complete MOA architecture remains an open empirical question. 7. FUTURE WORK The primary direction for future work is empirical validation. A minimal viable MOA pipeline could be constructed using: • Commercially available VOC sensor arrays • Portable spectrometry hardware • Open-source molecular sensing platforms A prototype system could follow this architecture: ┌─────────────────────────┐ │ Physical Environment │ └────────────┬────────────┘ │ ▼ ┌─────────────────────────┐ │ Molecular Sensor │ └────────────┬────────────┘ │ ▼ ┌─────────────────────────┐ │ VOC Detection │ └────────────┬────────────┘ │ ▼ ┌─────────────────────────┐ │ Compound Identification│ └────────────┬────────────┘ │ ▼ ┌─────────────────────────┐ │ Structured Molecular │ │ Representation │ └────────────┬────────────┘ │ ▼ ┌─────────────────────────┐ │ LLM │ └────────────┬────────────┘ │ ▼ ┌─────────────────────────┐ │ Olfactory Interpretation│ └─────────────────────────┘ 7.1 EXPERIMENTAL EVALUATION Controlled experiments could be conducted using a defined scent corpus containing known molecular compositions. Potential evaluation metrics could include: • Scent identification accuracy • Compound-to-scent reasoning accuracy • Robustness to concentration changes • Performance on unseen scent combinations • False-positive rate • Confidence calibration • Agreement with human olfactory assessments 7.2 RETRIEVAL-AUGMENTED GENERATION A second research direction would investigate whether Retrieval-Augmented Generation (RAG) over chemical databases can improve molecular interpretation. A retrieval system could provide the LLM with verified information about: • Molecular structures • Odor descriptors • Concentration thresholds • Known applications • Chemical properties • Known associations between compounds and odors This could reduce hallucination and improve factual grounding. 7.3 SPECIALIZED FINE-TUNING Another direction would be evaluating whether fine-tuning on specialized: • Olfactory chemistry literature • Experimental odor datasets • Molecular-to-odor mappings • Human olfactory assessments provides meaningful improvements over general-purpose LLMs. 7.4 HYBRID MOLECULAR ENCODERS Future research could investigate whether MOA requires an LLM at every stage of interpretation. A possible alternative would be a hybrid architecture: Molecular Sensor │ ▼ Molecular Encoder │ ▼ Chemical Representation │ ▼ LLM │ ▼ Semantic Reasoning │ ▼ Natural-Language Output 8. CONCLUSION This paper presented the Molecular Olfaction Architecture (MOA), a conceptual framework for extending LLM-based perception into the olfactory domain. The central proposal is that an AI system may not necessarily need to learn smell entirely from scratch. Instead: A molecular sensing layer could convert real-world odors into structured chemical information, while an LLM could use its existing chemical and semantic knowledge to interpret that information. The informal proof of concept presented in this paper suggests that the reasoning layer is technically plausible: an existing general-purpose LLM can receive a list of chemical compounds and generate a coherent interpretation of the associated olfactory profile. However, this demonstration does not validate the complete architecture. The major unresolved challenge is the sensory interface itself: Developing affordable, portable, reliable, and sufficiently precise molecular detection hardware capable of functioning as an olfactory encoder. In vision-language systems, cameras provide the sensory bridge between the physical world and the model. In audio-language systems, microphones serve a similar function. MOA proposes that molecular sensors could eventually play an analogous role for olfaction. If successful, this could transform smell from a purely descriptive concept that AI can talk about into a physical sensory modality that AI can actually measure, interpret, and reason about. ────────────────────────────────────────────────────────────────────────────────
## AUTHOR STATEMENT This document is a speculative concept paper and has not undergone peer review. No empirical end-to-end experiments were conducted. The author declares no conflicts of interest. No funding was received.
────────────────────────────────────────────────────────────────────────────────